Back

Molecular Plant

Elsevier BV

Preprints posted in the last 7 days, ranked by how well they match Molecular Plant's content profile, based on 39 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
Two evolutionary histories in one nucleus: genome remodeling and allelic regulation underlying heterosis in hybrid oil palm

Su, X.; Peng, Y.; Yang, X.; Zhang, F.; Xu, Q.; Ma, Z.; Dong, Y.; Zhou, L.; Xue, H.; Cao, X.; Zou, Z.; Wang, Y.; Zhou, Y.; Zeng, X.

2026-08-31 genomics 10.64898/2026.08.27.747553 medRxiv
Top 0.7%
1.7%
Show abstract

Oil palm (Elaeis) is the primary source of global vegetable oil. Interspecific hybrids of Elaeis exhibit pronounced heterosis by integrating two distinct subgenomes into a single nucleus, effectively combining the high yield of African oil palm (E. guineensis) with the high unsaturated fatty acid content and disease resistance of American oil palm (E. oleifera). However, the genetic basis underlying heterosis is still unclear. Here, we combine phased genome assembly, comparative genomics, evolutionary genomics and haplotype-aware transcriptomics to unravel the genetic architecture of heterosis of hybrid oil palm. We assemble the highly heterozygous F1 genome ('Reyou 40', 3.75% heterozygosity) into a complete 1.73 Gb T2T haplotype (HapG) and a 1.84 Gb near-T2T haplotype (HapO with17 gaps). Despite 91.56% sequence identity, HapG and HapO diverged in LTR-RT occurrence and PAV affected genes, showing complementary biases in lipid metabolism and stress responses, respectively. Evolutionary genomics revealed that ancient WGDs preserved the palm family. Whereas lineage-specific lipid-related gene expansions in oil palm. Six ancient introgressed regions (~64 Mb) in HapG were reshaped by transposable elements and tandem duplication, showing an enrichment of genes related to resistance and lipid metabolism. Transcriptomically, 82.2% of allelic gene pairs maintained balanced expression, accompanied by parental functional complementarity and dosage buffering, revealing a potential regulatory basis for coordinating parental genetic differences in the hybrid genome. These haplotype-resolved genomic resources offer vital targets for understanding heterosis and accelerating oil palm molecular breeding.

2
A subgenome-resolved and chromosome-scale reference genome assembly of allotetraploid wheat wild relative Aegilops peregrina

Singh, J.; Gudi, S.; Maughan, P. J.; Gill, U.; Gupta, R.

2026-08-30 genomics 10.64898/2026.08.28.747929 medRxiv
Top 1%
0.8%
Show abstract

Aegilops peregrina is a wild allotetraploid wheat wild relative and an important source of genetic diversity for stress tolerance and agronomic traits. Here, we report a subgenome-resolved, chromosome-scale reference genome assembly of a drought tolerant and stem rust resistant Ae. peregrina accession PI 604178 generated using PacBio HiFi and Hi-C sequencing. The 10.13 Gb assembly contains 98.81% of sequence anchored to 14 pseudomolecules representing the seven S and seven U chromosomes, with contig and scaffold N50 values of 25.84 and 746.48 Mb, respectively. The assembly achieved a consensus quality value of 74.61, 97.83% k-mers completeness, and 99.9% BUSCO completeness. LTR Assembly Index values of 20.43 and 18.79 for the S and U subgenomes, respectively, further supported high continuity across repeat-rich regions. Repetitive elements comprise 85.93% of chromosome-anchored assembly. We annotated 59,910 high-confidence protein-coding genes, with comparable gene representation across the two subgenomes. This reference genome provides a high-quality genomic framework for comparative analyses, characterization of important loci regulating agronomic and resilience related traits, and sequence-guided exploitation of Ae. peregrina allelic diversity for wheat improvement.

3
Molecular dissection of zinc-mediated immunity in Arabidopsis thaliana

Escudero, V.; Hoang, C. V.; Garcia-Molina, A.; De, A.; Armas, A. M.; Brueckner, D.; Ferreira Sanchez, D.; Bueschl, C.; Doppler, M.; van der Ent, A.; Schuhmacher, R.; Gonzalez-Guerrero, M.; Jorda, L.

2026-08-31 plant biology 10.64898/2026.08.28.747889 medRxiv
Top 1%
0.8%
Show abstract

Zinc is an essential micronutrient at low concentrations, yet it becomes toxic at slightly higher ones. This is exploited by plants as an effective defensive strategy. However, the molecular components that are involved zinc-mediated immunity remain poorly defined. Here, we show that mixed-linked {beta}-1,3/1,4-glucans naturally occurring in microbial and grass cell walls and used as an agrobiological solution, trigger zinc accumulation in the Arabidopsis apoplast and upregulate the expression of the zinc transporters HMA2 and HMA4. This response occurs independently of salicylic acid, jasmonic acid and ethylene-mediated signalling pathways, but it requires the LysM receptor kinases CERK1, LYK4 and LYK5, indicating a specific pattern triggered immunity-associated mechanism. We further demonstrate that hma2hma4 mutants display constitutive activation of a broad set of defence-related genes, yet this transcriptional reprogramming is insufficient to confer resistance against the necrotrophic fungus Plectosphaerella cucumerina BMM. Moreover, metabolomic profiling highlights the contribution of specialized metabolites to this defective defence output. Altogether, our findings reveal that zinc-mediated toxicity constitutes a defence mechanism integrated into the immune response triggered by specific microbial or damage associated molecular patterns.

4
MORC and MOM1 spatially constrain RNA Polymerase V chromatin positioning to shape DNA methylation landscapes

He, X.; Li, Z.; Xue, Y.; Guo, J.; Liu, X.; Feng, S.; Zhong, Z.; Jacobsen, S. E.

2026-08-31 plant biology 10.64898/2026.08.30.748062 medRxiv
Top 1%
0.5%
Show abstract

Plant-specific RNA Polymerase V (Pol V) transcribes noncoding RNAs in the RNA-directed DNA methylation pathway, thereby influencing gene expression and genome stability by controlling de novo DNA methylation. However, the mechanisms governing precise chromatin localization and transcriptional activities of Pol V remain elusive. Here we show that Pol V localization is spatially constrained by the chromatin regulators microrchidia (MORC) and MORPHEUS' MOLECULE 1 (MOM1). MORC and MOM1 promote Pol V occupancy at sites near active chromatin, whereas their loss leads to redistribution of Pol V into CMT3-enriched heterochromatin, accompanied by noncoding RNA transcription, small RNA production and DNA methylation. Our findings reveal a combinatorial model in which recruitment, spatial constraint and DNA methylation feedback collectively define Pol V chromatin distribution and epigenetic function.

5
Reference-guided comparative genomics of seven Indonesian rice cultivars identifies conserved gene space and trait-associated sequence candidates

Purwestri, Y. A.; Wicaksono, A.; Nurbaiti, S.; Purba, N. T.; Retnaningati, D.; Restiani, R.; Kumalasari, N.; Nuringtyas, T. R.; Handayani, V. D. S.

2026-08-29 genomics 10.64898/2026.08.26.747264 medRxiv
Top 1%
0.4%
Show abstract

Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92x to 41.58x, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.

6
The first chromosome-scale genome assembly of Blumeria graminis f. sp. avenae provides insights into genome evolution and host specialization

Ding, Y.; Zhang, P.; Ociepa, T.; Nucia, A.; Guan, H.; Kowalczyk, K.; Park, R. F.; Okon, S.

2026-08-30 genomics 10.64898/2026.08.28.747853 medRxiv
Top 2%
0.4%
Show abstract

Blumeria graminis f. sp. avenae (Bga), the causal agent of oat powdery mildew, is one of the most host-specialized members of the B. graminis species complex. Despite its agricultural importance, the lack of a high-quality reference genome has limited studies of host specialization, virulence evolution and comparative genomics in this pathogen. Here, we generated the first chromosome-scale genome assembly of Bga using an integrative approach combining long- and short-read sequencing, Hi-C scaffolding and transcriptome data. The Bga genome exhibits hallmark features of powdery mildew fungi, including extensive repeat content and low gene density. Comparative analyses revealed that genome expansion is primarily associated with historical transposable element proliferation rather than recent transpositional activity. Genome organization is consistent with a functionally stratified "one-speed" model, in which genes associated with pathogenicity, including predicted effectors and infection-responsive genes, are preferentially located in transposable element-rich regions characterized by reduced synteny conservation and extended intergenic spaces. In contrast, conserved genes are concentrated in compact genomic regions and maintain strong syntenic conservation across cereal-infecting formae speciales. Hi-C analyses demonstrated a highly structured chromatin architecture and revealed genome organization patterns associated with infection-related gene expression. Comparative genomic analyses indicated that host specialization in Bga is driven by localized diversification of a relatively small subset of genes rather than large-scale genome restructuring. These results provide the first high-quality genomic resource for Bga and offer new insights into the evolutionary mechanisms underlying host specialization in powdery mildew fungi.

7
Jasmonate-responsive group IX AP2/ERF transcription factors control the biosynthesis of benzylisoquinoline alkaloids

Yamada, Y.; Tatsumi, Y.; Inagaki, A.; Shitan, N.; Sato, F.

2026-08-31 plant biology 10.64898/2026.08.30.748054 medRxiv
Top 2%
0.2%
Show abstract

Although the biosynthetic pathways of benzylisoquinoline alkaloids (BIAs) have been extensively investigated in several plant species, their transcriptional regulatory mechanisms remain only partially understood. Jasmonate (JA)-responsive group IX APETALA2/Ethylene Responsive Factor (AP2/ERF) transcription factors (TFs) are well-known regulators of specialized plant metabolism, including the biosynthesis of various alkaloids. However, their specific roles in BIA biosynthesis remain largely elusive. Here, we isolated five novel group IX AP2/ERF TFs, designated Benzylisoquinoline alkaloid Jasmonate-responsive AP2/ERF (BJE1-5), from Coptis japonica. Phylogenetic analysis revealed that Benzylisoquinoline alkaloid Jasmonate-responsive AP2/ERF (BJE) proteins belong to subclades distinct from group IXa, which contains well-known AP2/ERF TFs involved in alkaloid biosynthesis. Transient expression analyses in C. japonica protoplasts demonstrated that certain BJEs, particularly CjBJE3 and CjBJE5, positively regulated BIA biosynthetic genes through a mutual regulatory network among BJE members. Moreover, CjBJE3 expression was regulated by CjbHLH1, a unique-type basic helix-loop-helix (bHLH) TF specific to BIA-producing plants. Furthermore, heterologous expression of CjBJE3 and CjBJE5 in cultured Eschscholzia californica cells significantly enhanced the overall BIA production, particularly by increasing end-product benzophenanthridine BIAs, highlighting several uncharacterized biosynthetic genes clustered in the genome. Our findings suggest that BIA-producing species have developed a specific regulatory network comprised of CjbHLH1 and BJE TFs, providing valuable clues for identifying novel biosynthetic enzymes.

8
PhenoStream: A Cyberinfrastructure for Automated and AI-Based Crop Trait Extraction from Aerial Imagery

Varela, S.; Ruhter, J.; Sacks, E.; Zheng, X.; Allen, D.; Hale, A.; Landry, C.; Kuang, X.; Long, B.; Zhu, Y.; Proma, S.; Kaur, S.; Jarquin, D.; Morrison, J.; Leakey, A.

2026-08-30 plant biology 10.64898/2026.08.26.747008 medRxiv
Top 3%
0.1%
Show abstract

The integration of digital technologies for high-throughput field phenotyping is critical for accelerating crop improvement in agriculture. However, extracting traits from remote sensing data remains constrained by fragmented workflows, manual intervention, and limited interoperability among existing tools, resulting in delays that hinder timely biological insight and decision-making. To address these challenges, we present PhenoStream (Phenotyping Streaming), a scalable, end-to-end cyberinfrastructure designed to automate the full lifecycle of aerial imagery-based phenotyping, from data acquisition to plot- and genotype-level inference. The framework integrates automated data ingestion from distributed field sites, geospatial processing, and AI-enabled trait extraction within a unified, user-accessible graphical interface. Its modular and extensible architecture supports adaptable trait modeling and seamless integration of new data sources, enabling deployment across diverse crops, environments, and experimental designs. We demonstrate the system across a large multi-location field trial network of bioenergy crops, where it enables high-throughput characterization of spatiotemporal growth dynamics, genotype-by-environment (GxE) interactions, and predictive modeling of key agronomic traits. By significantly reducing processing latency and manual effort, the platform facilitates near-real-time analysis and reproducible workflows. This work establishes a generalizable and scalable pathway for operationalizing very-high-spatial resolution aerial phenotyping in agricultural research. By bridging data acquisition and analytics, the end-to-end cyberinfrastructure provides a foundation for integrating heterogeneous and unstructured data streams--including remote sensing, environmental, and management data--toward data-driven decision making in agriculture.

9
Eucalyptus microRNA Archive (EMA): a multi-study and cross-condition curated database of microRNAs in Eucalyptus grandis

Aires Teixeira, J. V.; Motta Venancio, T.; Quintanilha-Peixoto, G.; Pimenta de Oliveira, K. K.

2026-08-31 plant biology 10.64898/2026.08.29.747619 medRxiv
Top 3%
0.1%
Show abstract

MicroRNAs (miRNAs) are key post-transcriptional regulators of development, stress response, and secondary cell wall formation in woody plants, yet annotations for Eucalyptus grandis, the world's most widely planted hardwood, remain fragmented across studies using incompatible discovery pipelines and filtering criteria. Here we present the Eucalyptus MicroRNA Archive (EMA), a curated, locus-resolved database integrating three independent small RNA sequencing datasets spanning vegetative tissue, somatic embryogenesis, and mechanically induced tension wood formation. Applying annotation criteria aligned with current plant miRNA standards, EMA catalogs 99 curated miRNAs (31 previously described, 68 novel) organized into 34 family-level groupings under a three-tier confidence system, known-reference-supported, multi-study replicated, or single-study, that preserves study-of-origin and sample-level evidence for every entry. Cross-study comparison showed that only 9 of 99 entries (9.1%) were independently supported by all three datasets, supporting an evidence-tiered rather than binary annotation scheme. Target prediction against the E. grandis transcriptome yielded 1,773 miRNA-target interactions spanning 764 loci, integrated into a combined miRNA-target and protein-protein interaction network. This network resolved into functionally coherent, mutually isolated clusters, including an miR482-associated NBS-LRR/TIR disease-resistance hub with a substantial translational-repression component, alongside modules enriched for ribosome biogenesis and translation, DNA replication, and nitrogen and carbohydrate metabolism. EMA is publicly accessible through an interactive web dashboard, with all curated data, source code, and analysis scripts openly available, providing a reproducible, extensible framework for E. grandis miRNA research and a template for similarly structured resources in other non-model woody species.

10
Isoprenoid Binding and Substrate Channeling in Drimenol Synthase, a Bifunctional Class II Terpene Cyclase-Phosphatase

Osika, K. R.; Leffler, M. E.; Czarnecki, B. A. R.; Christianson, D. W.

2026-08-31 biochemistry 10.64898/2026.08.29.748031 medRxiv
Top 3%
0.1%
Show abstract

More than one thousand bifunctional terpene synthases combining prenyltransferase and terpene cyclase activities have been identified in bacteria and fungi, but only a handful of enzymes have been identified that combine terpene cyclase activity with a downstream processing activity. Drimenol synthase from the marine bacterium Aquimarina spongiae (AsDMS) consists of a class II terpene cyclase that converts farnesyl diphosphate into drimenyl diphosphate, and a haloacid dehalogenase-like phosphatase that hydrolyzes drimenyl diphosphate to generate the sesquiterpene alcohol drimenol. The first crystal structure of AsDMS to be reported revealed the architecture of domain assembly as well as dimeric quaternary structure, establishing a structural chemical foundation for cyclization and hydrolysis mechanisms [K. R. Osika, M. N. Gaynes, D. W. Christianson (2025) Proc. Natl. Acad. Sci. U.S.A. 122, e2506584122]. Here, we report crystal structures of the catalytically-inactive double mutant, D33A-D323A AsDMS, complexed with farnesyl diphosphate, geranyl diphosphate, and dimethylallyl diphosphate, which bind in the active sites of both the cyclase and phosphatase domains. Molecular recognition of the diphosphate group dominates binding interactions in both active sites. In the cyclase active site, only farnesyl diphosphate is sufficiently long for its terminal isoprenoid C=C bond to bind adjacent to the catalytic general acid that would initiate the cyclization cascade in the wild-type enzyme. In the phosphatase active site, all isoprenoid diphosphate groups bind similarly, but isoprenoid chain conformations vary. These structures provide a foundation for understanding substrate recognition and catalysis in both active sites. Finally, we present kinetic evidence suggesting that substrate channeling is operative in wild-type AsDMS.

11
Visual LLM-guided consensus spatial domain detection with L-STAR

Zhao, C.; Ji, Z.

2026-08-29 bioinformatics 10.64898/2026.08.25.747158 medRxiv
Top 3%
0.1%
Show abstract

Spatial domain detection is a central task in spatial transcriptomics, yet existing methods exhibit highly variable performance across datasets. We introduce L-STAR, a visual LLM-guided, consensus-based framework that leverages the visual reasoning capacity of large language models to adaptively rank and integrate spatial domain detection methods. L-STAR achieves robust and consistently improved performance, outperforming single spatial domain detection methods across diverse datasets.

12
Yeast Dhx29 promotes translation progression by unwinding structured mRNA in the ribosomal A-site

Chitoiu, L.; Denk, T.; Müller, M. B. D.; Berninghausen, O.; Becker, T.; Thoms, M.; Beckmann, R.

2026-08-31 biochemistry 10.64898/2026.08.24.746666 medRxiv
Top 3%
0.1%
Show abstract

mRNAs can form stable structures that need to be resolved to facilitate translation. During translation initiation in mammals, the scanning 48S complex requires the helicase activity of DHX29 to unwind stable mRNA structures that cannot be resolved by eIF4A. Here, we show that the yeast DHX29 homolog, Ylr419w (Dhx29), has a similar function during translation on elongating 80S ribosomes. Cryo-EM analyses show that the Dhx29 helicase module is positioned at the mRNA entry channel to engage mRNA, while its double-stranded RNA-binding domain (dsRBD) senses hairpin-forming mRNA in the ribosomal A-site. By selective ribosome profiling, we observed that Dhx29 is associated with transcripts that form RNA structures, such as stable tetraloops. Dhx29 mutants with perturbed helicase activity enrich 80S with hairpins in the A-site, as well as ribosome collisions, while a mutant lacking the N-terminal dsRBD sensor domain loses the specificity for such ribosomes. We thus propose that Dhx29 functions in translation elongation by resolving structured mRNA formed in the ribosomal A-site through its 3'-5' helicase activity and pulling on the mRNA from its 3' end.

13
Vipsania: Unsupervised Deep Gene Finding

Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.

2026-08-30 bioinformatics 10.64898/2026.08.26.747235 medRxiv
Top 3%
0.1%
Show abstract

Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.

14
Utilising nuclear encoded plastid DNA to identify donors of grass-to-grass lateral gene transfer

Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.

2026-08-29 evolutionary biology 10.64898/2026.08.26.747220 medRxiv
Top 4%
0.1%
Show abstract

Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.

15
S-Palmitoylation stabilizes OGT and the OGT-PPP1CC complex

Lu, X.; Xu, T.; Li, J.; Liu, Y.; Zhou, W.; Wang, K.; Niu, C.; Tang, N.; Zhang, L.; Li, J.

2026-08-31 biochemistry 10.64898/2026.08.29.747956 medRxiv
Top 4%
0.1%
Show abstract

O-linked {beta}-N-acetylglucosamine (O-GlcNAc) transferase (OGT) is the sole writer for intracellular O-GlcNAcylation. It catalyzes O-GlcNAcylation of thousands of protein substrates, but relatively less is known about the post-translational modifications that occur on OGT itself. Herein, we demonstrate that OGT is S-palmitoylated at Cys-472 and Cys-477, which is mediated by the S-acyltransferase Zinc Finger DHHC-Type Palmitoyl transferase 14 (zDHHC14) and removed by acyl protein thioesterase 2 (APT2). S-Palmitoylation stabilizes OGT by shunting it away from the lysosomal chaperone-mediated autophagy (CMA) pathway, as S-palmitoylation decreases the interaction between OGT and heat shock cognate 70 kDa protein (HSC70), the CMA chaperone. Via label-free quantitative mass spectrometry, we find that S- palmitoylation elevates the affinity between OGT and protein phosphatase 1 catalytic subunit gamma (PPP1CC), but not PPP1CB. We further demonstrate that S-palmitoylation of OGT augments binding with Yes-associated protein-1 (YAP), a protein that associates with PPP1CC, and subsequently enhances YAP O-GlcNAcylation. Our work unearths S-palmitoylation of OGT and CMA-mediated degradation of lysosomal OGT, the orchestration of which finetunes the activity of key OGT complexes, such as OGT-PPP1CC, and contributes to OGT substrate selectivity.

16
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 4%
0.0%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

17
Predicting COVID-19 hospitalisation and common disease risk from comorbid diagnoses in 13 million individuals

Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,

2026-09-01 health informatics 10.64898/2026.08.27.26361302 medRxiv
Top 4%
0.0%
Show abstract

Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.

18
PCGS: biomarker and risk group identification for Pediatric Cancers via explainable Graph neural networks with Shapley values

Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.

2026-09-01 health informatics 10.64898/2026.08.27.26361540 medRxiv
Top 4%
0.0%
Show abstract

Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.

19
Vaginal Steaming Practices Among Women At Mulago Sti Clinic, Uganda: Prevalence, Lactobacilli Proportion Differences, And Associated Factors.

Kakai, D.; Twinamasiko, N.; Kigozi, E.; Namutale, R.; Mutesi, B. A.; Bagaya, J.; Akinyi, L.; Kajumbula, H.; Nakubulwa, S.

2026-09-01 obstetrics and gynecology 10.64898/2026.08.28.26361407 medRxiv
Top 4%
0.0%
Show abstract

Abstract Background: Vaginal steaming has gained popularity among women for reasons known best to them. However, the effects of vaginal steaming on vaginal lactobacilli levels remain poorly understood and understudied. This study investigated the prevalence, assessed differences in the presence of bacterial vaginosis (BV) among women who practiced vaginal steaming and those who do not and determined the factors associated with vaginal steaming among women attending Mulago Sexually Transmitted Infections (STI) clinic in Uganda. Methods: This study utilized a cross-sectional design to enroll 181 women aged 18 to 49 years who were systematically sampled at Mulago STI clinic. Interviews were conducted to obtain the demographic characteristics of the participants. Vaginal swabbing and Gram staining were done to attain lactobacilli counts by microscopy which were categorized using the Nugent score. Data were analyzed using Stata. The potential confounding effects of other variables on the relation between vaginal steaming and the presence of bacterial vaginosis as well as factors associated with vaginal steaming were assessed using modified Poisson regression. Results: Prevalence of vaginal steaming was 40.3%, (95% confidence interval (CI) 33.0% - 48.0%). There were 41.1% women who practiced vaginal steaming occasionally, 53.4% who used hot water having herbs and 78.1% who practiced vaginal steaming for medical reasons. There was no difference in the presence of bacterial vaginosis when women who practiced vaginal steaming were compared to those who did not (p-value = 0.286). Factors that were significantly associated with vaginal steaming included having experienced vaginal issues (aPR = 0.07, 95% CI 0.01 - 0.12, p value = < 0.001) and contraceptive use (aPR= 0.52, 95% CI 0.37 - 0.72, p value = 0.001). Conclusions: About 2 in every 5 women at Mulago STI clinic reported to have indulged in vaginal steaming. There was no difference in the presence of bacterial vaginosis when women who practiced vaginal steaming were compared to those who did not. Having experienced vaginal issues and contraceptive use were significantly associated with vaginal steaming among women at Mulago STI clinic, Uganda. The Ministry of Health of Uganda should establish targeted screening and treatment for bacterial vaginosis alongside other sexually transmitted infections.

20
Dynamic Clinical States and Transitions During the First 72 Hours of Intensive Care After Acute Stroke

LEI, P.; XU, Y.; ZHANG, Y.

2026-09-01 intensive care and critical care medicine 10.64898/2026.08.30.26361738 medRxiv
Top 4%
0.0%
Show abstract

Background: The condition of a patient with acute stroke often changes within hours of ICU admission. Prognostic work here targets fixed endpoints predicted from admission data, and trajectory phenotyping assigns one label per patient. We used longitudinal ICU data to identify interpretable dynamic clinical states, characterize transitions between them, and relate the current state to later events. Methods: Retrospective cohort study of 6368 adults with acute stroke in MIMIC IV v3.1. The first 72 h were divided into twelve 6-hour windows, and a hidden Markov model was fitted to 21 neurological, physiological and organ support variables. State number was chosen against criteria fixed before fitting: statistical fit, restart stability, state occupancy and clinical interpretability. Generalized estimating equations related the current state to new mechanical ventilation and vasopressor use within 12 h, and to ICU death within 72 h. Eleven sensitivity analyses assessed the robustness of the state solution. Results: Four states were selected: neurologically preserved-low support, neurological impairment low support, impairment renal dysfunction and impairment-respiratory support (63.3%, 7.8%, 11.8% and 17.1% of windows). Within 72 h, 40.3% of patients changed state at least once, and transitions ran in both directions rather than along a single severity gradient. States were identified without outcome data, yet ICU mortality by last state ranged from 2.9% to 43.9%. Adjusted for age, sex, subtype and Charlson index, the current state remained associated with organ-support escalation and death. State prevalence differed by at most 1.1 percentage points between training and test sets, and 10 of 11 sensitivity analyses gave a stable four-state solution (ARI 0.754 0.955). Conclusions: The early ICU course of acute stroke can be represented as movement among a small number of clinically interpretable states. The representation was reproducible in a held out set and across admission eras, but requires validation in an independent database before any clinical use.